Wearable devices measure body signals such as skin conductance, pulse, skin temperature and movement, and machine-learning models can use these signals to detect acute stress. Recent work has added explainability so that a user can see which signal influenced a decision. However, most studies use only a single explanation method, judge that explanation visually, and rarely report how confident the model is. This raises a practical question: can these explanations be trusted? This paper studies that question on the WESAD dataset under strict subject-independent (Leave-One-Subject-Out) evaluation. Five deep models (LSTM, BiLSTM, GRU, CNN-LSTM, Transformer) and five classical models are trained on identical folds. Explanations are generated with three methods (Integrated Gradients, SHAP, LIME) and their agreement is measured across six held-out participants; explanations are further examined with faithfulness tests, cross-architecture comparison, counterfactual and temporal analysis; and predictive uncertainty is estimated with deep ensembles and assessed for calibration. Three findings are reported. First, the five deep architectures perform almost identically (0.846–0.863 accuracy) and are statistically indistinguishable (all Holm-corrected p = 1.00, small effect sizes), while a classical Random Forest is more accurate (0.911); after Holm–Bonferroni correction no individual classical-vs-deep pair is significant at 15 subjects, but the direction is consistent across every comparison, with medium-to-large effect sizes for nine of the ten comparisons (rank-biserial 0.50–0.68) and a smaller effect for the tenth (Logistic Regression vs Transformer, +0.25), so at this data scale increasing architectural complexity is not the productive direction. Second, across six subjects the three explanation methods agree only partially, heart rate is the single most important channel in two to four of six subjects depending on the method, and pairwise method agreement is low-to-moderate and highly variable (Spearman ? from 0.10 ± 0.45 to 0.70 ± 0.20), which shows that a single explanation method is insufficient on its own and that reliability must be measured rather than assumed; faithfulness is in the expected direction (mean insertion AUC 0.794 > deletion AUC 0.756) but modest and subject-dependent. Third, the ensemble is well calibrated (Expected Calibration Error 0.065) and, under a coverage–risk (deferral) analysis, answering only the most-confident 50% of cases raises selective accuracy from 0.884 to 0.960. The study concludes that trustworthiness, measured explanation reliability together with calibration, rather than accuracy alone, is the appropriate objective for wearable stress detection.
Introduction
The text presents a trustworthiness assessment framework for wearable-based stress detection using machine learning and deep learning. Rather than focusing only on how accurately stress can be detected, the study asks whether the model's predictions, explanations, and confidence levels can be trusted, especially in health-related applications.
Stress causes measurable physiological changes such as increased heart rate, increased sweating, and changes in skin temperature. Wearable devices can continuously capture these signals, making them useful for automatic stress detection. However, traditional questionnaires are subjective and cannot provide continuous monitoring. Although many machine-learning and deep-learning systems achieve high accuracy, two major problems remain: lack of explainability and unreliable confidence.
The study addresses these problems through four main contributions:
Model comparison: Five deep-learning models and five classical machine-learning models are evaluated using the same subject-independent validation procedure. The study examines whether model architecture significantly affects performance.
Explanation agreement: Three explainable-AI methods—Integrated Gradients (IG), SHAP, and LIME—are applied to the same model. Their explanations are quantitatively compared across multiple participants rather than relying on a single visual explanation.
Explanation verification: Explanations are tested using deletion/insertion faithfulness, cross-architecture consistency, counterfactual analysis, and temporal occlusion to determine whether highlighted physiological signals genuinely influence predictions.
Uncertainty and calibration: The study evaluates whether model confidence corresponds to actual accuracy and uses coverage–risk analysis to allow the model to abstain from uncertain predictions.
Dataset and preprocessing
The research uses the WESAD (Wearable Stress and Affect Detection) dataset, containing data from 15 participants recorded with an Empatica E4 wristband. Five physiological channels are analyzed:
Electrodermal activity (EDA)
Blood volume pulse (BVP)
Skin temperature (TEMP)
Acceleration (ACC)
Heart rate (HR), derived from BVP
The signals are resampled to 4 Hz and divided into 60-second windows with a 5-second stride. Only windows containing at least 90% of a single condition are retained, resulting in 6,198 labelled windows, with approximately 30% stress and 70% non-stress samples.
Importantly, the study uses Leave-One-Subject-Out (LOSO) cross-validation. This prevents information from the same person appearing in both training and testing data, which could artificially inflate performance. Because the windows overlap heavily, the researchers correctly treat the 15 participants—not the individual windows—as the independent statistical units.
Models used
The study compares ten models:
Deep-learning models:
LSTM
BiLSTM
GRU
CNN-LSTM
Transformer with self-attention
Classical machine-learning models:
Logistic Regression
Random Forest
Support Vector Machine (SVM)
k-Nearest Neighbors (k-NN)
Histogram Gradient Boosting
The classical models use 75 engineered features derived from the five physiological channels, including statistical and frequency-domain features.
Explainability and trustworthiness
The research separates four important concepts:
Explanation agreement: Whether IG, SHAP, and LIME identify similar important physiological channels.
Faithfulness: Whether the features highlighted by an explanation actually affect the model's prediction.
Calibration: Whether predicted confidence accurately reflects the probability of being correct.
Overall trustworthiness: Whether agreement, faithfulness, calibration, and consistency with physiological knowledge collectively support confidence in the model.
For example, the study reports strong agreement between some explanation methods, including IG and LIME with a Spearman correlation of 1.00 for an illustrative participant, while IG and SHAP show an average correlation of 0.70.
Main research gap
Previous stress-detection studies often focus on achieving high prediction accuracy but have three limitations:
Explanations are often generated using only one method and may be checked visually rather than quantitatively.
Model uncertainty and calibration are rarely reported.
Some studies use validation procedures that allow the same participant's data to appear in both training and testing, potentially producing overly optimistic results.
This study attempts to address all three issues.
Conclusion
This paper examined whether explanations produced by wearable stress-detection models can be trusted. Using WESAD under strict LOSO, five deep and five classical models were compared with multiple-comparison correction and effect sizes; three explanation methods were applied to the same model and compared across six held-out participants; explanations were verified through faithfulness testing and cross-architecture analysis; and predictive uncertainty was estimated, calibrated and analysed through a coverage–risk curve. The study finds that the deep architectures are statistically indistinguishable, that a classical model is at least as strong at this data scale, that explanation methods agree only partially and with high variance across subjects, so explanation reliability must be measured, not assumed, and that a calibrated ensemble with deferral raises selective accuracy from 0.884 to 0.960 at 50% coverage. The central conclusion is that trustworthiness, measured explanation reliability together with calibration, rather than accuracy alone, is the appropriate objective for this field.
Limitations and future work. The dataset has only 15 participants and was recorded in a controlled laboratory; the task was binary; the multi-subject explainability analysis used six participants and the illustrative faithfulness figures come from one; and, because windows overlap, the effective number of independent observations is the number of subjects, which limits statistical power. Attribution results can also depend on the Integrated Gradients baseline. Future work will extend the evaluation to additional datasets and to all subjects, examine multi-level stress classification, incorporate heart-rate-variability features, investigate per-subject personalisation and calibration, and assess on-device feasibility.
References
[1] M. K. Moser, M. Ehrhart, and B. Resch, “An Explainable Deep Learning Approach for Stress Detection in Wearable Sensor Measurements,” Sensors, vol. 24, no. 16, 5085, 2024.
[2] P. Schmidt, A. Reiss, R. Duerichen, C. Marberger, and K. Van Laerhoven, “Introducing WESAD, a Multimodal Dataset for Wearable Stress and Affect Detection,” Proc. 20th ACM ICMI, 2018, pp. 400–408.
[3] M. Kyrou, I. Kompatsiaris, and P. C. Petrantonakis, “Deep Learning Approaches for Stress Detection: A Survey,” IEEE Trans. Affective Computing, vol. 16, no. 2, pp. 499–517, 2025.
[4] M. Sundararajan, A. Taly, and Q. Yan, “Axiomatic Attribution for Deep Networks,” Proc. 34th ICML, 2017, pp. 3319–3328.
[5] S. M. Lundberg and S.-I. Lee, “A Unified Approach to Interpreting Model Predictions,” NeurIPS 30, 2017, pp. 4765–4774.
[6] M. T. Ribeiro, S. Singh, and C. Guestrin, “Why Should I Trust You? Explaining the Predictions of Any Classifier,” Proc. 22nd ACM SIGKDD, 2016, pp. 1135–1144.
[7] B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles,” NeurIPS 30, 2017, pp. 6402–6413.
[8] Y. Gal and Z. Ghahramani, “Dropout as a Bayesian Approximation,” Proc. 33rd ICML, 2016, pp. 1050–1059.
[9] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On Calibration of Modern Neural Networks,” Proc. 34th ICML, 2017, pp. 1321–1330.
[10] S. Hochreiter and J. Schmidhuber, “Long Short-Term Memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997.
[11] A. Vaswani et al., “Attention Is All You Need,” NeurIPS 30, 2017, pp. 5998–6008.
[12] S. Gedam and S. Paul, “A Review on Mental Stress Detection Using Wearable Sensors and Machine Learning Techniques,” IEEE Access, vol. 9, pp. 84045–84066, 2021.
[13] G. Giannakakis et al., “Review on Psychological Stress Detection Using Biosignals,” IEEE Trans. Affective Computing, vol. 13, no. 1, pp. 440–460, 2022.
[14] J. DeYoung et al., “ERASER: A Benchmark to Evaluate Rationalized NLP Models,” Proc. 58th ACL, 2020, pp. 4443–4458.
[15] P. Bobade and M. Vani, “Stress Detection with Machine Learning and Deep Learning Using Multimodal Physiological Data,” Proc. 2nd ICIRCA, 2020, pp. 51–57.
[16] Y. S. Can, B. Arnrich, and C. Ersoy, “Stress Detection in Daily Life Scenarios Using Smartphones and Wearable Sensors: A Survey,” J. Biomedical Informatics, vol. 92, 103139, 2019.